Skip to content

[Config] [AMD][DSV4] [tune][gfx950] wo_b/wq_b a8w8 blockscale bpreshuffle configs - #5279

Merged
coderfeli merged 5 commits into
ROCm:mainfrom
karverma-amd:karverma/dsv4-gfx950-gemm-tune
Sep 7, 2026
Merged

coderfeli merged 5 commits into
ROCm:mainfrom
karverma-amd:karverma/dsv4-gfx950-gemm-tune

Conversation

@karverma-amd

Copy link
Copy Markdown
Contributor

Summary

Add MI355X (gfx950) tuned a8w8 block-scale bpreshuffle GEMM configs for two
DeepSeek-V4-Pro attention projections that had no gfx950 rows and were falling back
to aiter's default config:

projection shape (N, K)
wo_b (o_proj up) 7168 × 1024
wq_b (q up-proj) 8192 × 1536

Tuned across the M ladder spanning concurrency 4–256 (decode M=1 … prefill M=16384).
The tuner selects tuned cktile at large M.

Kernel result (MI355X, gfx950)

shape M default tuned speedup
wo_b (7168×1024) 16384 224 µs 148 µs ~1.34×
wq_b (8192×1536) 16384 ~277 µs 224 µs ~1.24×

End-to-end result (DeepSeek-V4-Pro, TP8, 128k ISL / 1k OSL)

Long-context prefill is where these projections run at large M most often (a 128k prompt
chunks into 16 prefill forwards at M=8192). TTFT, 3 reps, conc 8:

rep default tuned
1 1682.9 ms 1549.4 ms
2 1531.1 ms 1520.3 ms
3 1956.5 ms 1844.0 ms
mean 1724 ms 1638 ms (−5.0%)

Tuned is faster in every rep (TTFT and total throughput). Decode is unaffected
(decode M is tiny), and GSM8K accuracy is unchanged (0.948).

How the config was generated

# untuned shapes: wo_b (7168,1024) + wq_b (8192,1536) across M = 1,2,4,...,16384
python3 csrc/ck_gemm_a8w8_blockscale/gemm_a8w8_blockscale_tune.py \
    --preshuffle --libtype all \
    -i aiter/configs/a8w8_blockscale_bpreshuffle_untuned_gemm.csv \
    -o aiter/configs/model_configs/dsv4_a8w8_blockscale_bpreshuffle_tuned_gemm.csv

How it was validated

Server (DeepSeek-V4-Pro, MI355X, TP8):

export SGLANG_USE_AITER=1 SGLANG_DEFAULT_THINKING=1 SGLANG_DSV4_REASONING_EFFORT=high
export SGLANG_USE_ROCM700A=0 SGLANG_HACK_FLASHMLA_BACKEND=unified_kv_triton
export AITER_BF16_FP8_MOE_BOUND=0 SGLANG_OPT_USE_AITER_BATCHED_GEMM=1
export SGLANG_OPT_UNIFIED_CACHE_FREE_OUT_OF_WINDOW_SLOTS=1 SGLANG_TIMEOUT_KEEP_ALIVE=900

python3 -m sglang.launch_server \
    --model-path deepseek-ai/DeepSeek-V4-Pro --tensor-parallel-size 8 \
    --attention-backend dsv4 --page-size 256 --swa-full-tokens-ratio 0.10 \
    --kv-cache-dtype fp8_e4m3 --enforce-shared-experts-fusion \
    --tool-call-parser deepseekv4 --reasoning-parser deepseek-v4 \
    --chunked-prefill-size 8192 --mem-fraction-static 0.89 \
    --max-running-requests 16 --cuda-graph-max-bs 16 --context-length 200000 \
    --speculative-algorithm EAGLE --speculative-num-steps 3 \
    --speculative-eagle-topk 1 --speculative-num-draft-tokens 4 --watchdog-timeout 3600

Accuracy — GSM8K (1319 questions, 1319 parallel):

python3 benchmark/gsm8k/bench_sglang.py --num-questions 1319 --parallel 1319 --port 8888
# -> Accuracy: 0.948  (unchanged vs default config)

Perf — 128k ISL / 1k OSL, streaming TTFT probe (conc 8, 3 reps):

# streaming client: NUM_PROMPTS concurrent prompts of ~INPUT_WORDS words
# (~1.06 tok/word => ~128k tokens), OUTPUT_LEN new tokens, reports TTFT/ITL/throughput
PORT=8888 NUM_PROMPTS=8 CONCURRENCY=8 OUTPUT_LEN=1024 INPUT_WORDS=120800 \
    python3 load_probe.py     # 3 reps; TTFT means in the table above

# equivalent with the stock sglang serving benchmark:
python3 -m sglang.bench_serving --backend sglang --port 8888 \
    --dataset-name random --random-input-len 131072 --random-output-len 1024 \
    --num-prompts 8 --max-concurrency 8

Notes / scope

  • gfx950 only; other architectures are unaffected (no rows added for them).
  • Additive change to dsv4_a8w8_blockscale_bpreshuffle_tuned_gemm.csv only; no existing
    rows modified.
  • The win is specific to large-M (prefill / long-context); it is intentionally
    neutral for small-M decode where these projections are a negligible fraction of the step.

@karverma-amd
karverma-amd requested a review from a team September 4, 2026 20:59
@github-actions

github-actions Bot commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

🏷️ CI Guide

Runs automatically on every PR:

  • ✅ Pre-checks (submodule verification, code formatting)
  • ✅ Aiter op tests (gfx942 + gfx950)
  • ✅ Triton tests on MI35X (only when aiter/ops/triton/** or related paths are changed)

Extended tests (opt-in via labels):

Label Tests
ci:gfx1250-ffm-triton Run the five-shard gfx1250 FFM Triton test suite
ci:triton-300x Run an additional Triton test job on MI300X in PRs; main branch always runs both MI35X and MI300X
multigpu Aiter multi-GPU tests on the 8-GPU runner
ci:sglang SGLang integration tests: DeepSeek-R1-MXFP4 accuracy, Qwen 3.5 accuracy
ci:atom ATOM benchmark: DeepSeek-R1-0528, GPT-OSS-120B
ci:atom_full ATOM accuracy suite for PR and main models from ATOM models_accuracy.json
ci:vllm vLLM benchmark: GPT-OSS-120B, DeepSeek-R1-0528, Kimi-K2.5
ci:all All standard extended tests (excludes ci:atom_full)

Only add ci:atom_full for FlyDSL or Triton upgrades.
Add labels via the sidebar or gh pr edit 5279 --add-label <label>

PR title tags & labels:
Component tags ([Triton/Gluon], [HIP], [CK], [ASM], ...) are added to the PR title and as PR labels automatically from the changed files and re-synced on every push — change-type tags like [fix]/[Perf], op tags like [MLA], and human labels (ci:*) are left untouched. Add the no-auto-title label to opt this PR out.

@github-actions github-actions Bot changed the title [tune][gfx950] DSv4 wo_b/wq_b a8w8 blockscale bpreshuffle configs [Config] [tune][gfx950] DSv4 wo_b/wq_b a8w8 blockscale bpreshuffle configs Sep 4, 2026
@github-actions github-actions Bot added the Config label Sep 4, 2026
@karverma-amd karverma-amd changed the title [Config] [tune][gfx950] DSv4 wo_b/wq_b a8w8 blockscale bpreshuffle configs [AMD][DSV4] [tune][gfx950] wo_b/wq_b a8w8 blockscale bpreshuffle configs Sep 4, 2026
@karverma-amd
karverma-amd force-pushed the karverma/dsv4-gfx950-gemm-tune branch from 6398c51 to 7496e7b Compare September 4, 2026 21:02
@github-actions github-actions Bot changed the title [AMD][DSV4] [tune][gfx950] wo_b/wq_b a8w8 blockscale bpreshuffle configs [Config] [AMD][DSV4] [tune][gfx950] wo_b/wq_b a8w8 blockscale bpreshuffle configs Sep 4, 2026
wo_b (o_proj up, N=7168 K=1024) and wq_b (q up-proj, N=8192 K=1536) had no
gfx950 tuned rows and fell back to aiter's default config on MI355X. Add tuned
rows across the M ladder spanning concurrency 4-256 (decode M=1 .. prefill
M=16384); the tuner selects tuned cktile at large M (wo_b 224->148us @m=16384,
~34% vs default).

Validated on DeepSeek-V4-Pro (MI355X, TP8): +5% TTFT at 128k/1k long-context;
GSM8K accuracy unchanged (0.948).

Co-authored-by: Cursor <cursoragent@cursor.com>
@karverma-amd
karverma-amd force-pushed the karverma/dsv4-gfx950-gemm-tune branch from 7496e7b to feb952f Compare September 4, 2026 21:10
karverma-amd and others added 2 commits September 5, 2026 15:54
wq_b previously had only M=4,8 in the small/mid range before jumping to 1024,
so decode/low-concurrency shapes fell back to aiter's default config. Add 20
autotuned rows (M=1,2,16,32,48,64,80,96,112,128,144,160,176,192,208,224,240,
256,384,512), bringing wq_b to parity with wo_b's ladder. ck wins at small M
(~6us), asm BpreShuffle tiles at mid M; all errRatio=0.0.

Tuned with: gemm_a8w8_blockscale_tune.py --preshuffle --libtype all (gfx950, MI355X).

Co-authored-by: Cursor <cursoragent@cursor.com>
…le config

The wq_b shape is comprehensively tuned for gfx950 in
a8w8_blockscale_bpreshuffle_tuned_gemm_dsv3.csv (64 rows), which merges into the
same runtime lookup (dedup key gfx,M,N,K). The dsv4 wq_b rows were therefore
cross-file duplicates and tripped the prebuild merge/dedup check
(RuntimeError: duplicate shape entries). dsv3's wq_b timings are within noise of
this tune, so removing them loses nothing. Keep only wo_b (7168x1024), which is
genuinely uncovered by any blockscale-family config. Net PR is now +15 wo_b rows;
merged set has 0 duplicate shapes.

Co-authored-by: Cursor <cursoragent@cursor.com>
HaiShaw
HaiShaw previously approved these changes Sep 6, 2026
@coderfeli

Copy link
Copy Markdown
Collaborator

add 32k config or <32k will go into default

coderfeli and others added 2 commits September 6, 2026 15:17
Reviewers (coderfeli, HaiShaw) noted the wo_b ladder stopped at M=16384, so
prefill shapes above 16384 fell back to aiter's default config. Add the
autotuned M=32768 row: cktile a8w8_blockscale_cktile_192x256x128_4x2x1_
16x16x128_intrawave_0x1x0_1, 284.5us, 1690 TFLOP/s, errRatio=0.0 (gfx950,
MI355X, --preshuffle --libtype all). Same kernel family as the M=16384 winner,
~linear scaling. wo_b ladder now 1..32768; no duplicate shapes in the
blockscale_bpreshuffle merge set.

Co-authored-by: Cursor <cursoragent@cursor.com>
@coderfeli
coderfeli merged commit 24a62b1 into ROCm:main Sep 7, 2026
58 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants